Papers with Knowledge-based visual question answering

3 papers
Hypergraph Transformer: Weakly-Supervised Multi-hop Reasoning for Knowledge-based Visual Question Answering (2022.acl-long)

Copied to clipboard

Challenge: Existing knowledge-based visual question answering tasks require weak supervision and no visual knowledge.
Approach: They propose a model which encodes high-level semantics of a question and a knowledge base and learns high order associations between them.
Outcome: The proposed model encodes high-level semantics of a question and a knowledge base, and learns high order associations between them.
Modality-Aware Integration with Large Language Models for Knowledge-Based Visual Question Answering (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to integrate multimodal knowledge in a modality-agnostic manner can be sub-optimal.
Approach: They propose a modality-aware integration with large language models (LLMs) that leverages multimodal knowledge for both image understanding and knowledge reasoning.
Outcome: The proposed model is able to bridge a tight inter-modal exchange while preserving insightful intra-modal learning.
Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge-based visual question answering systems rely on Concept-Net and Wikipedia to obtain external knowledge.
Approach: They propose a visual retriever-reader pipeline that uses a natural language knowledge base and a Visual retriever to retrieve relevant knowledge.
Outcome: The proposed method significantly improves the visual retriever-reader pipeline on the OK-VQA benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations